Integrating Adaptive Beam-forming and Auditory Features for Robust Large Vocabulary Speech Recognition

نویسندگان

Xie Sun

Peter Li

Manli Zhu

Qiru Zhou

چکیده

We demonstrate a system to integrate adaptive beam-forming and auditory features in order to improve speech recognition accuracy in noisy environments. Adaptive beam-forming based on a microphone array can utilize spatial information to improve the sound recording signal-to-noise ratio (SNR) on a focused speaker for robust speech recognition. Auditory features based on modeling the signal processing functions in the hearing system have shown to largely improve speech recognition accuracy under noisy conditions. According to our experiments, when both adaptive beam-forming and the auditory features are integrated, an absolute gain of more than 50% over a baseline on speech recognition accuracy is achieved when 5dB white noise is added.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

An Information-Theoretic Discussion of Convolutional Bottleneck Features for Robust Speech Recognition

Convolutional Neural Networks (CNNs) have been shown their performance in speech recognition systems for extracting features, and also acoustic modeling. In addition, CNNs have been used for robust speech recognition and competitive results have been reported. Convolutive Bottleneck Network (CBN) is a kind of CNNs which has a bottleneck layer among its fully connected layers. The bottleneck fea...

متن کامل

Mutual Information Based Dynamic Integration of Multiple Feature Streams for Robust Real-Time LVCSR

We present a novel method of integrating the likelihoods of multiple feature streams, representing different acoustic aspects, for robust speech recognition. The integration algorithm dynamically calculates a frame-wise stream weight so that a higher weight is given to a stream that is robust to a variety of noisy environments or speaking styles. Such a robust stream is expected to show discrim...

متن کامل

Regularized minimum variance distortionless response-based cepstral features for robust continuous speech recognition

In this paper, we present robust feature extractors that incorporate a regularized minimum variance distortionless response (RMVDR) spectrum estimator instead of the discrete Fourier transform-based direct spectrum estimator, used in many front-ends including the conventional MFCC, to estimate the speech power spectrum. Direct spectrum estimators, e.g., single tapered periodogram, have high var...

متن کامل

Feature extraction based on hearing system signal processing for robust large vocabulary speech recognition

A new auditory-based feature extraction algorithm for robust speech recognition is developed from modeling the signal processing functions in the hearing system. Usually, the performance of acoustic models trained in clean speech drops significantly when tested on noisy speech; thus recognition systems cannot work robustly in the field even when they have good performance in labs. To address th...

متن کامل

شبکه عصبی پیچشی با پنجره‌های قابل تطبیق برای بازشناسی گفتار

Although, speech recognition systems are widely used and their accuracies are continuously increased, there is a considerable performance gap between their accuracies and human recognition ability. This is partially due to high speaker variations in speech signal. Deep neural networks are among the best tools for acoustic modeling. Recently, using hybrid deep neural network and hidden Markov mo...

متن کامل

ذخیره در منابع من

ذخیره در منابع من قبلا به منابع من ذحیره شده

{@ msg_add @}

با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره شماره

صفحات -

تاریخ انتشار 2012

Integrating Adaptive Beam-forming and Auditory Features for Robust Large Vocabulary Speech Recognition

نویسندگان

چکیده

منابع مشابه

An Information-Theoretic Discussion of Convolutional Bottleneck Features for Robust Speech Recognition

Mutual Information Based Dynamic Integration of Multiple Feature Streams for Robust Real-Time LVCSR

Regularized minimum variance distortionless response-based cepstral features for robust continuous speech recognition

Feature extraction based on hearing system signal processing for robust large vocabulary speech recognition

شبکه عصبی پیچشی با پنجره‌های قابل تطبیق برای بازشناسی گفتار

عنوان ژورنال:

اشتراک گذاری